Papers with German language

18 papers
Neural OCR Post-Hoc Correction of Historical Corpora (2021.tacl-1)

Copied to clipboard

Challenge: Optical character recognition (OCR) is crucial for a deeper access to historical collections.
Approach: They propose a neural approach based on a combination of recurrent (RNN) and deep convolutional network (ConvNet) to correct OCR transcription errors.
Outcome: The proposed model reduces the word error rate of 32.3% by more than 89% on a historical book corpus in German language.
PunKtuator: A Multilingual Punctuation Restoration System for Spoken and Written Text (2021.eacl-demos)

Copied to clipboard

Challenge: Prior punctuation restoration methods have focused on using lexical features, prosodic features or combination of both.
Approach: They propose a multitask modeling approach to restore punctuation in multiple high resource languages using acoustic models and a computational model.
Outcome: The proposed system can restore punctuation in Germanic, Romanic and low resource languages without extensive knowledge of grammar or syntax.
A Corpus for Argumentative Writing Support in German (2020.coling-main)

Copied to clipboard

Challenge: In today's world most information is readily available. Consequently, the sole reproduction of information is losing attention.
Approach: They propose an annotation approach to capture claims and premises of arguments and their relations in student-written peer reviews on business models in german language.
Outcome: The proposed annotation scheme guides annotators to moderate agreement with the proposed scheme on 50 persuasive student-written peer reviews on business models.
Robustness Evaluation of the German Extractive Question Answering Task (2025.coling-main)

Copied to clipboard

Challenge: Existing evaluation benchmarks for Question Answering systems only include EM and F1 scores, but they overlook critical factors for the deployment of QA systems.
Approach: They propose to define an evaluation method specifically tailored to the German language to evaluate the robustness of German QA models.
Outcome: The proposed method extends existing methods to German language . it shows that all models are vulnerable to character-level perturbations .
Acquiring a Formality-Informed Lexical Resource for Style Analysis (2021.eacl-main)

Copied to clipboard

Challenge: lexico-statistics analysis of formality levels in written communication has long been dominated by application concerns, such as authorship and plagiarism assignment problems.
Approach: They propose a lexicon with entries ordered by their degree of (in)formality and let crowdworkers assess the enlarged set of lexical items on a continuous informal-formal scale as a gold standard for evaluation.
Outcome: The proposed lexicon is evaluated on a German-language email corpus and is then evaluated by crowdworkers.
Merkel Podcast Corpus: A Multimodal Dataset Compiled from 16 Years of Angela Merkel’s Weekly Video Podcasts (2022.lrec-1)

Copied to clipboard

Challenge: a dataset of 16 years of (almost) weekly Internet podcasts of former german chancellor Angela Merkel is presented.
Approach: They propose to curate a German podcast corpus from 16 years of podcasts of former german chancellor Angela Merkel using audio-visual-text methods.
Outcome: The proposed pipeline can be used to curate other datasets of similar nature, such as talk show contents.
Supporting Land Reuse of Former Open Pit Mining Sites using Text Classification and Active Learning (2021.acl-long)

Copied to clipboard

Challenge: open pit mines left many regions worldwide inhospitable or uninhabitable . aforementioned information has to be acquired to ensure safety and validity of land reuse .
Approach: They propose a workflow for supporting the post-mining management of former open pit mines in the eastern part of Germany . they use active learning to perform multi-label sentence classification for two categories of restrictions and seven categories of topics .
Outcome: The proposed system supports the post-mining management of former lignite open pit mines in the eastern part of Germany.
GRhOOT: Ontology of Rhetorical Figures in German (2022.lrec-1)

Copied to clipboard

Challenge: GRhOOT is a domain ontology of rhetorical figures in the German language . the goal is to allow for easier detection of non-literal language based tasks .
Approach: GRhOOT is a domain ontology of 110 rhetorical figures in the german language . the goal is to allow for easier detection and sentiment analysis .
Outcome: The ontology of rhetorical figures in the German language is based on 110 rhetorical figure domains . the goal is to make the ontologies more accurate and to allow for easier detection .
SuperGLEBer: German Language Understanding Evaluation Benchmark (2024.naacl-long)

Copied to clipboard

Challenge: a new set of German-pretrained models are being released, but no established, diverse and systematic evaluation suite is available for them.
Approach: They assemble a Natural Language Understanding benchmark suite for the German language and evaluate 10 existing German-pretrained models.
Outcome: The proposed benchmark suite evaluates 10 German-pretrained models on 29 tasks . the results show that encoder models are good choices for most tasks, but not all .
A Joint Approach to Compound Splitting and Idiomatic Compound Detection (2020.lrec-1)

Copied to clipboard

Challenge: Compounding is a common word-formation process in Germanic languages . high productivity and low corpus frequency of compounds increase vocabulary size .
Approach: They develop a deep learning-based approach to noun compound splitting and idiomatic compound detection for the German language.
Outcome: The proposed approach outperforms the current state of the art in noun compound splitting and idiomatic compound detection for the German language.
Modeling Persuasive Discourse to Adaptively Support Students’ Argumentative Writing (2022.acl-long)

Copied to clipboard

Challenge: Argumentation is an omnipresent rudiment of daily communication and thinking . humans struggle to develop argumentation skills due to a lack of individual and instant feedback in their learning process.
Approach: They propose an argumentation annotation approach to model argumentative discourse in student-written business model pitches and embed it into an adaptive writing support system for students that provides individual argumentation feedback.
Outcome: The proposed method annotates a corpus of 200 business model pitches in german and measures their self-efficacy and ease-of-use in a real-world writing exercise.
Abstract Text Summarization: A Low Resource Challenge (D19-1)

Copied to clipboard

Challenge: Existing datasets for multilingual text summarization are difficult to construct and lack of human knowledge and language processing abilities in computers makes text summaries a challenging task.
Approach: They propose an iterative data augmentation approach which uses synthetic data along with the real summarization data for the German language.
Outcome: The proposed system improves on the development and test sets on the German language text using the state-of-the-art “Transformer” model.
FinCorpus-DE10k: A Corpus for the German Financial Domain (2024.lrec-main)

Copied to clipboard

Challenge: a predominantly German corpus of financial documents is available for the first time . financial text is characterized by a unique vocabulary with implications including sentiment analysis .
Approach: They propose a predominantly German financial corpus comprising 12.5k PDF documents . they hope it will fill this gap and foster further research in the financial domain .
Outcome: The proposed corpus is the first non-email German financial corpus available . it aims to provide insights into financial discourse in the German language and multilingually.
German SRL: Corpus Construction and Model Training (2024.lrec-main)

Copied to clipboard

Challenge: Existing semantic role annotation resources are lacking for German.
Approach: They propose a translation-based approach to train German semantic role models using semantic annotations and alignment models.
Outcome: The proposed method achieves competitive evaluation scores, but avoids limitations of previous approaches.
From Witch’s Shot to Music Making Bones - Resources for Medical Laymen to Technical Language and Vice Versa (2020.lrec-1)

Copied to clipboard

Challenge: Information we share online unveils directly or indirectly information about our lifestyle and health situation.
Approach: They propose a dataset which annotates medical laymen and technical expressions in a patient forum and a set of medical synonyms and definitions.
Outcome: The proposed dataset annotates medical laymen and technical expressions in a patient forum along with a set of medical synonyms and definitions.
Summarization Corpora of Wikipedia Articles (2020.lrec-1)

Copied to clipboard

Challenge: Using Wikipedia articles, we extract summarization data for other languages.
Approach: They propose a process to extract Wikipedia summarization corpora and apply it to the German language.
Outcome: The proposed method can be applied to the German language and compares to baselines.
PopAut: An Annotated Corpus for Populism Detection in Austrian News Comments (2024.lrec-main)

Copied to clipboard

Challenge: Populism is a phenomenon that is noticeably present in political landscapes worldwide . prior work on populism analysis focused on analyzing populist content expressed by politicians .
Approach: They present a corpus of news comments annotated for populism in the german language . they use machine learning to detect populist comments in text .
Outcome: The proposed corpus outperforms existing dictionaries for populism detection in text . it features 1,200 comments collected between 2019-2021 .
Using Pre-Trained Language Models in an End-to-End Pipeline for Antithesis Detection (2024.lrec-main)

Copied to clipboard

Challenge: Rhetorical figures are a "departure from the normal usage" of language . features of metaphors, irony and sarcasm enhance performance of several NLP tasks.
Approach: They propose a pipeline approach to detect rhetorical figures using large language models by splitting text into phrases and identifying parallel phrases with a syntactically parallel structure.
Outcome: The proposed approach outperforms state-of-the-art methods by an F1 score of 65.11 %.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations